Papers with diagnostic evaluation of state-of-the-art or near-state-of-the-art
RuBia: A Russian Language Bias Detection Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data. |
| Approach: | They present a bias detection dataset specifically designed for the Russian language, dubbed RuBia, which is divided into 4 domains: gender, nationality, socio-economic status, and diverse. |
| Outcome: | The proposed dataset is designed to detect bias in the Russian language and is based on 2,000 unique sentence pairs spread over 19 subdomains. |